Papers with dialog systems

57 papers
Optimizing NLU Reranking Using Entity Resolution Signals in Multi-domain Dialog Systems (2021.naacl-industry)

Copied to clipboard

Challenge: In dialog systems, the Natural Language Understanding component makes the interpretation decision before the mentioned entities are resolved.
Approach: They propose to leverage Entity Resolution (ER) features in NLU reranking to learn model weights . they propose a score distribution matching method to ensure the models are calibrated .
Outcome: The proposed approach outperforms the baseline model on multiple domain evaluations.
Designing, Evaluating, and Learning from Humans Interacting with NLP Models (2023.emnlp-tutorial)

Copied to clipboard

Challenge: This tutorial will cover how to conduct human-in-the-loop usability evaluations to ensure that models are capable of interacting with humans.
Approach: They will provide a systematic overview of key considerations and effective approaches for studying human-NLP model interactions.
Outcome: This tutorial will cover how to conduct human-in-the-loop usability evaluations to ensure that models are capable of interacting with humans.
Do Neural Dialog Systems Use the Conversation History Effectively? An Empirical Study (P19-1)

Copied to clipboard

Challenge: Neural generative models are becoming more popular when building conversational agents.
Approach: They propose to study the sensitivity of neural dialog models to unnatural perturbations . they experiment with 10 different types of perturbations on 4 multi-turn dialog datasets .
Outcome: The proposed model is sensitive to unnatural changes or perturbations on 4 multi-turn dialog datasets.
Semantic Diversity for Natural Language Understanding Evaluation in Dialog Systems (2020.coling-industry)

Copied to clipboard

Challenge: a dialog system is used to evaluate NLU models using aggregated metrics on a large number of utterances.
Approach: They propose a method to generate a test set with high semantic diversity for NLU evaluation in dialog systems.
Outcome: The proposed test sets are based on high diversity of utterances from different regions of the utteration embedding space.
NoEl: An Annotated Corpus for Noun Ellipsis in English (2020.lrec-1)

Copied to clipboard

Challenge: Ellipsis resolution is an important step to improve the accuracy of mainstream natural language processing tasks such as information retrieval, event extraction, dialog systems, etc.
Approach: They extend the study of ellipsis by annotating a corpus for noun ellippsis and closely related phenomenon using the first hundred movies of Cornell Movie Dialogs Dataset.
Outcome: The proposed corpus has 946 instances of exophoric and endophorical noun ellipsis, making it the biggest resource of nouns in English, to the best of our knowledge.
ChatEval: A Tool for Chatbot Evaluation (N19-4)

Copied to clipboard

Challenge: open-domain dialog systems are difficult to evaluate due to lack of standardization and standardization in evaluation procedures.
Approach: They propose a framework for human evaluation of chatbots that augments existing tools . researchers can submit their trained models to the ChatEval web interface . reproducibility and model assessment for opendomain dialog systems is challenging .
Outcome: The proposed framework provides a web-based hub for researchers to compare their models with baselines and prior work.
DIAGRAPH: An Open-Source Graphic Interface for Dialog Flow Design (2023.acl-demo)

Copied to clipboard

Challenge: Dialog systems have gained attention as a convenient way for users to access information in a more personalized manner.
Approach: They present a graphical dialog flow editor built on ADVISER toolkit . it provides a clean and intuitive graphical interface for creating dialog systems .
Outcome: The tool is based on the ADVISER toolkit and is evaluated with subject-experts . it is able to quickly prototype dialog systems and provide a test bed for students learning about dialog systems.
Are the Tools up to the Task? an Evaluation of Commercial Dialog Tools in Developing Conversational Enterprise-grade Dialog Systems (N19-2)

Copied to clipboard

Challenge: Existing toolsets are incomplete in meeting the goal of building effective dialog systems, authors say .
Approach: They compare dialog tools available from a number of companies to determine their strengths and weaknesses . they provide quantitative and qualitative results in three main areas: natural language understanding, dialog, and text generation .
Outcome: The toolsets are incomplete, but they are compared to other tools to determine their strengths and weaknesses.
Code-Switching Can be Better Aligners: Advancing Cross-Lingual SLU through Representation-Level and Prediction-Level Alignment (2024.acl-short)

Copied to clipboard

Challenge: Existing code-switching-based cross-lingual spoken language understanding frameworks are limited to low-resource languages.
Approach: They propose a cross-lingual spoken language understanding framework that leverages both code-switched and original sentences to achieve multi-level alignment.
Outcome: The proposed framework can achieve multi-level alignment on two benchmarks across ten languages.
Estimating Soft Labels for Out-of-Domain Intent Detection (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods to detect out-of-dominance (OOD) intents are limited by the lack of OOD samples.
Approach: They propose an adaptive soft pseudo labeling method that can estimate soft labels for pseudo OOD samples when training OOD detectors.
Outcome: The proposed method outperforms competing methods on three benchmark datasets and consistently outperformed previous methods.
Intent Features for Rich Natural Language Understanding (2021.naacl-industry)

Copied to clipboard

Challenge: generic dialog systems, or chatbots, are increasingly popular, but most industrial dialog systems are built for specific clients and use cases.
Approach: They propose a new neural network architecture that allows for domain and topic agnostic properties of intents that can be learnt from syntactic cues only.
Outcome: The proposed model improves on baselines for identifying intent features in a deployed, multi-intent natural language understanding module.
Usnea: An Authorship Tool for Interactive Fiction using Retrieval Based Semantic Parsing (2020.acl-demos)

Copied to clipboard

Challenge: Interactive Fiction is a genre of art and entertainment that is well known in the context of video games . authorship tools for IF define some structure of a story and provide suggested algorithms or software itself to realize this structure in a form that a reader can digest .
Approach: They propose to use retrieval based semantic parsing to create a novel type of Interactive Fiction.
Outcome: The proposed novel type of Interactive Fiction is open-source and features novel algorithms and models based on the IF community's existing models.
OodGAN: Generative Adversarial Network for Out-of-Domain Data Generation (2021.naacl-industry)

Copied to clipboard

Challenge: Existing models for OOD detection work with text, but they do not work directly with the text.
Approach: They propose to use a sequential generative adversarial network (SeqGAN) based model to generate OOD data for a given domain automatically.
Outcome: The proposed model outperforms state-of-the-art in OOD detection metrics for ROSTD and OSQ datasets.
Towards Large-Scale Interpretable Knowledge Graph Reasoning for Dialogue Systems (2022.findings-acl)

Copied to clipboard

Challenge: Existing systems that require extensive labor to process user requests are limited in their reasoning capabilities and require extensive manual effort to design.
Approach: They propose a method that allows a transformer model to walk on a large-scale knowledge graph to generate responses by reasoning over differentiable knowledge graphs.
Outcome: The proposed method allows a transformer model to walk on a large-scale knowledge graph to generate responses.
Outlier Detection for Improved Data Quality and Diversity in Dialog Systems (N19-1)

Copied to clipboard

Challenge: Existing methods to detect outliers in text have been neglected in NLP . outlier detection is a problem in dialog systems where text is often no more than a few sentences in length.
Approach: They propose a technique that uses sentence embeddings to detect outliers in short texts using neural sentence embeds and distance-based outlier detection.
Outcome: The proposed technique detects outliers in a corpus of short texts while generating highly diverse corpora that produce more robust intent classification and slot-filling models.
Construction and Analysis of a Multimodal Chat-talk Corpus for Dialog Systems Considering Interpersonal Closeness (2020.lrec-1)

Copied to clipboard

Challenge: a large-scale multimodal dialog corpus is needed to accelerate research on dialog systems that can handle social signals and verbal information.
Approach: They construct a multimodal dialog corpus focusing on the relationship between speakers and 19 pairs of participants.
Outcome: The proposed system is based on a multimodal dialog corpus of 19,303 utterances (10 hours) from 19 pairs of participants.
Learning Low-Resource End-To-End Goal-Oriented Dialog for Fast and Reliable System Deployment (2020.acl-main)

Copied to clipboard

Challenge: Existing end-to-end dialog systems perform less effectively when data is scarce.
Approach: They propose a Meta-Dialog System which combines meta-learning and human-machine collaboration to improve dialog learning by a new extended-bAbI dataset and a transformed MultiWOZ dataset.
Outcome: The proposed system outperforms non-meta-learning baselines on a new extended-bAbI dataset and a transformed MultiWOZ dataset for low-resource goal-oriented dialog learning.
DeepPavlov Dream: Platform for Building Generative AI Assistants (2023.acl-demo)

Copied to clipboard

Challenge: open-source DeepPavlov Dream Platform is designed for development of complex dialog systems . platform supports modular approach to implementation of conversational agents .
Approach: open-source DeepPavlov Dream Platform is designed for development of complex dialog systems . platform includes a conversational orchestrator called DeepPvlov Agent to coordinate asynchronous dialog pipeline .
Outcome: The open-source DeepPavlov Dream Platform is designed for development of complex dialog systems like Generative AI Assistants.
Novel Feature Discovery for Task-Oriented Dialog Systems (2023.findings-eacl)

Copied to clipboard

Challenge: Prior work on novelty detection limits the scope of features represented by novel single intents to those represented by multiple user-perceived fine-grained features belonging to the same intent.
Approach: They propose to use a feature discovery technique to discover novel features from user utterances rather than single intent discovery to classify them into slots.
Outcome: The proposed approach consistently detects novel features from user utterances on two datasets.
Where to Go for the Holidays: Towards Mixed-Type Dialogs for Clarification of User Goals (2022.acl-long)

Copied to clipboard

Challenge: a dialog system posits that users have figured out clear and specific goals . but in many real-world scenarios, users struggle to figure out specific goals by determining all the necessary slots.
Approach: They propose a mixed-type dialog model with a Prompt-based continual learning mechanism . they collect 5k dialog sessions and 168k utterances for 4 dialog types and 5 domains .
Outcome: The proposed model provides user-goal-related knowledge to help figure out clear and specific goals . it can be extended to any specific type by utilizing existing dialog corpora effectively.
DiaBiz – an Annotated Corpus of Polish Call Center Dialogs (2022.lrec-1)

Copied to clipboard

Challenge: DiaBiz is a large corpus of phone conversations from different business domains . it contains nearly 410 hours of recordings and over 3 million words of transcribed speech.
Approach: They introduce DiaBiz, a large, annotated, multimodal corpus of Polish telephone conversations . it is a multimodal, multi-modal corpor of 4036 phone conversations from nine different domains .
Outcome: The corpus of 4036 phone conversations in Poland is 410 hours long and contains over 3 million words of transcribed speech.
Out-of-domain Detection based on Generative Adversarial Network (D18-1)

Copied to clipboard

Challenge: Existing methods for out-of-domain (OOD) detection require huge effort to collect OOD sentences.
Approach: They propose to use only in-domain (IND) sentences to build a generative adversarial network (GAN) of which the discriminator generates low scores for OOD sentences.
Outcome: The proposed method is most accurate compared to existing methods on multi-domain dialog systems.
Database Search Results Disambiguation for Task-Oriented Dialog Systems (2022.naacl-main)

Copied to clipboard

Challenge: Task-oriented dialog systems can't handle multiplesearch results when querying a database due to the lack of such scenarios in existing datasets.
Approach: They propose a task that focuses on disambiguating database search results by synthetically generating turns through a pre-defined grammar and collecting human paraphrases for a subset.
Outcome: The proposed task improves performance on DSR-disambiguation even in the absence of in-domain data, suggesting it can be learned as a universal dialog skill.
Zero-shot Generalization in Dialog State Tracking through Generative Question Answering (2021.eacl-main)

Copied to clipboard

Challenge: Existing methods for Dialog State Tracking do not generalize well to new domains and unseen slots.
Approach: They propose an ontology-free framework that queries for unseen constraints and slots in multi-domain task-oriented dialogs using a conditional language model pre-trained on substantive English sentences.
Outcome: The proposed framework improves goal accuracy in zero-shot domain adaptation settings by up to 9% over the previous state-of-the-art on the MultiWOZ 2.1 dataset.
Unsupervised Discrete Sentence Representation Learning for Interpretable Neural Dialog Generation (P18-1)

Copied to clipboard

Challenge: Existing encoder-decoder dialog models cannot output interpretable actions as in traditional systems.
Approach: They propose an unsupervised discrete sentence representation learning method that integrates with existing encoder-decoder dialog models for interpretable response generation.
Outcome: The proposed model can be integrated with existing encoder-decoder dialog models and discover interpretable semantics via either auto encoding or context predicting.
CASA-NLU: Context-Aware Self-Attentive Natural Language Understanding for Task-Oriented Chatbots (D19-1)

Copied to clipboard

Challenge: Prior work on contextual NLU has been limited in terms of the types of contextual signals used and the understanding of their impact on the model.
Approach: They propose a context-aware self-attentive NLU model that uses multiple signals over a variable context window, such as previous intents, slots, dialog acts and utterances, in addition to the current user uttered.
Outcome: The proposed model outperforms a baseline model on two conversational datasets yielding a gain of up to 7% on the IC task.
Sentiment Adaptive End-to-End Dialog Systems (P18-1)

Copied to clipboard

Challenge: Existing methods to train dialog systems only consider semantic inputs and under-utilize other user information.
Approach: They propose to include user sentiment in the end-to-end learning framework to make dialog systems more user-adaptive and effective.
Outcome: The proposed system improves on a bus information search task with sentiment information.
TIAGE: A Benchmark for Topic-Shift Aware Dialog Modeling (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing dialog models can generate on-topic utterances but struggle to proactively switch topics.
Approach: They propose a topic-shift aware dialog benchmark based on human topic shift annotations.
Outcome: The proposed benchmark enables chatbots to generate topic-shift responses while still struggling to decide when to change topic.
Manipulating the Perceived Personality Traits of Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Psychology research has long explored aspects of human personality like extroversion, agreeableness and emotional stability, three of the personality traits that make up the ‘Big Five’.
Approach: They propose to use text generated from large language models to evaluate perceived personality traits and to frame them as tools for controlling personas in dialog systems.
Outcome: The proposed models predict personality traits in different contexts and can be manipulated in a predictable way.
Dependency Parsing for Spoken Dialog Systems (D19-1)

Copied to clipboard

Challenge: Dependency parsing of conversational input can help to understand dialogs . currently available annotation schemes do not adapt well to spoken human-machine dialogs.
Approach: They propose an annotation scheme that extends Universal Dependencies guidelines to spoken dialogs.
Outcome: The proposed scheme disambiguates relationships between entities extracted from dialogs . it is better than existing models on public datasets and fine-tuned on ConvBank data .
Unsupervised Dialog Structure Learning (N19-1)

Copied to clipboard

Challenge: Current dialog systems require human experts to design the dialog structure, which is time consuming and sometimes insufficient to satisfy various customer needs.
Approach: They propose to extract dialog structure using a modified VRNN model with discrete latent vectors.
Outcome: The proposed model outperforms existing models on the ability to predict unseen data and is faster and more effective in a reinforcement learning setting.
A Semi-Supervised Stable Variational Network for Promoting Replier-Consistency in Dialogue Generation (D19-1)

Copied to clipboard

Challenge: Existing methods favor uninformative and non replier-specific responses due to lack of relevant information guidance.
Approach: They propose to use a semi-supervised variable network to generate replier-specific responses . they use vMF as latent space to obtain stable KL performance .
Outcome: The proposed model outperforms baseline models on two large conversation datasets and generates diverse and replier-specific responses.
Robots-Dont-Cry: Understanding Falsely Anthropomorphic Utterances in Dialog Systems (2022.emnlp-main)

Copied to clipboard

Challenge: Dialog systems often output human-like responses, but some are impossible for a machine to say.
Approach: They collect ratings on the feasibility of 900 two-turn dialogs from 9 data sources . they build classifiers and explore how modeling configuration might affect output permissibly .
Outcome: The proposed model can be used to train human-like dialogs, but it is not anthropomorphic.
Towards Emotional Support Dialog Systems (2021.acl-long)

Copied to clipboard

Challenge: Emotional support is a crucial ability for many conversation scenarios, including social interactions, mental health support, and customer service chats.
Approach: They propose an Emotional Support Conversation task and an ESC Framework to train emotional support into dialog systems.
Outcome: The proposed framework provides an example of an Emotional Support Conversation task and shows that it is more effective than existing models.
Gated Mechanism Enhanced Multi-Task Learning for Dialog Routing (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for dialog routing are mostly heuristic and cannot achieve high-quality performance.
Approach: They propose a multi-task learning framework with a dialog encoder and two tailored gated mechanism modules to solve this problem.
Outcome: The proposed model can play the role of hierarchical information filtering and is non-invasive to existing dialog systems.
Frugal Prompting for Dialog Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are used in natural language processing tasks with an unrealistic speed and effectiveness.
Approach: They propose more compact ways of providing dialog history information while ensuring good performance and reducing model’s inference-API costs.
Outcome: The proposed models have the optimal usable-information density while maintaining good performance and reducing model’s inference-API costs.
Rethinking Coherence Modeling: Synthetic vs. Downstream Tasks (2021.eacl-main)

Copied to clipboard

Challenge: Coherence models are typically evaluated only on synthetic tasks, which may not be representative of their performance in downstream applications.
Approach: They compare models' performance on synthetic sentences with those on retrieval-based dialog.
Outcome: The proposed models perform poorly on synthetic sentences and retrieval-based dialog tasks.
Zero-shot User Intent Detection via Capsule Neural Networks (D18-1)

Copied to clipboard

Challenge: Existing methods to classify intents are labor-intensive and time-consuming as intents will be diverse and new intents may be involved.
Approach: They propose a zero-shot intent detection problem which aims to detect emerging user intents where no labeled utterances are currently available.
Outcome: The proposed model can discriminate emerging intents when no labeled utterances are available in training data.
Conversational Grounding: Annotation and Analysis of Grounding Acts and Grounding Units (2024.lrec-main)

Copied to clipboard

Challenge: Successful conversations often rest on common understanding, says a researcher . despite recent advances in dialog systems, there is a noticeable deficit in their grounding capabilities .
Approach: They propose to use a framework to build conversational grounding in dialogs . they propose to analyze two dialog corpora using grounding acts and grounding units .
Outcome: The proposed model shows that language models are not enough to ground dialogs with machines . the proposed model can be used to test the performance of existing Language Models .
Generating Responses with a Specific Emotion in Dialog (P19-1)

Copied to clipboard

Challenge: EmoDS can express emotions in both ways, but it is difficult to scale to large datasets.
Approach: They propose an emotional dialog system that can express emotions in both ways . they use strong emotional words and neutral words to increase the intensity of emotions .
Outcome: The proposed system performs better than baselines in BLEU, diversity and quality of emotional expression.
Personalized Transformer for Explainable Recommendation (2021.acl-long)

Copied to clipboard

Challenge: Recent years have witnessed the successful application of natural language generation.
Approach: They propose a model that uses user and item IDs to predict the words in the target explanation to make personalized Transformer.
Outcome: The proposed model outperforms BERT on the explainable recommendation task in terms of effectiveness and efficiency.
Domain-Aware Dependency Parsing for Questions (2021.findings-acl)

Copied to clipboard

Challenge: Pre-trained parsers perform poorly on domain-specific questions, a paper argues . retraining parser with domain- specific questions is expensive, as these require linguistic expertise.
Approach: They propose an automatic labeled domain question generation framework leveraging domain knowledge and seed domain questions.
Outcome: The proposed framework improves state-of-the-art parsers on domain questions.
Learning End-to-End Goal-Oriented Dialog with Multiple Answers (D18-1)

Copied to clipboard

Challenge: Existing methods for dialog learning assume there is only one correct next utterance . a significant drop in performance is seen in existing methods for evaluating dialog systems .
Approach: They propose a method that assumes there is only one correct next utterance in a dialog . they propose bAbI dialog tasks that introduce valid next .
Outcome: The proposed method improves performance and achieves 47.3% accuracy on permuted-bAbI dialog tasks.
Unsupervised Natural Language Generation with Denoising Autoencoders (D18-1)

Copied to clipboard

Challenge: Unsupervised approaches to generating text from structured data are costly to obtain and limited to a limited domain.
Approach: They propose an unsupervised approach that learns its parameters without the slot pairs on target sequences only.
Outcome: The proposed approach can generate sentences out of corrupted data without supervision . it can be used in question answering and dialog systems, the authors show .
Automatically Select Emotion for Response via Personality-affected Emotion Transition (2021.findings-acl)

Copied to clipboard

Challenge: Existing studies focus on rendering specified emotions in responses, yet the individual difference in emotion expression is overlooked.
Approach: They propose to equip a dialog system with personality and enable it to select emotions in responses like humans.
Outcome: The proposed system can select emotions in responses like humans by simulating the emotion transition of humans in conversation.
Hierarchical Transformer for Task Oriented Dialog Systems (2021.naacl-main)

Copied to clipboard

Challenge: Existing models for dialog generation are challenging to train using the standard Seq2Seq models.
Approach: They propose a framework for Hierarchical Transformer Encoders that can be morphed into any hierarchical transformer by using specially designed attention masks and positional encodings.
Outcome: The proposed framework can be morphed into any hierarchical encoder, including HRED and HIBERT like models, by using specially designed attention masks and positional encodings.
Taskmaster-1: Toward a Realistic and Diverse Dialog Dataset (D19-1)

Copied to clipboard

Challenge: a lack of high quality conversational data is limiting progress in dialog systems . we present a dataset of 13,215 task-based dialogs .
Approach: They propose a task-based dialog dataset which includes 13,215 task-related dialogs . they use a two-person, spoken "Wizard of Oz" approach and a "self-dialog" approach .
Outcome: The taskmaster-1 dataset contains 13,215 task-based dialogs comprising six domains.
P4: Plug-and-Play Discrete Prompting for Large Language Models Personalization (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit impressive capabilities in following instructions, but manually prompting them to exhibit certain personalities may result in sub-optimal performance.
Approach: They propose a plug-and-play prompting method to manipulate Large Language Models with distinct human-like personality traits by appending discrete personalized suffixes to query or dialog histories and focusing exclusively on influential tokens.
Outcome: The proposed method outperforms other prompting methods and model editing methods on four models ranging from 1.1B to 13B and achieves 79.9% accuracy in customizing LLMs’ personalities.
The R-U-A-Robot Dataset: Helping Avoid Chatbot Deception by Detecting User Questions About Human or Non-Human Identity (2021.acl-long)

Copied to clipboard

Challenge: We analyze 2,500 phrasings related to the intent of “Are you a robot?” and 2,500 adversarially selected utterances to determine whether systems are non-human.
Approach: They analyze 2,500 phrasings related to the intent of "Are you a robot?" and 2,500 adversarially selected utterances to determine whether systems are non-human.
Outcome: The proposed model and two systems fail to confirm non-human intent, and the proposed model is complex.
Towards Automatic Evaluation of Dialog Systems: A Model-Free Off-Policy Evaluation Approach (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for evaluation of dialog systems are expensive and not scalable . a framework for estimating human evaluation scores is proposed to bridge this gap .
Approach: They propose a framework for estimating human evaluation scores based on off-policy evaluation . they use language quality metrics for single-turn response generation given a fixed context .
Outcome: The proposed framework outperforms existing methods in terms of correlation with human evaluation scores.
“I’d rather just go to bed”: Understanding Indirect Answers (2020.emnlp-main)

Copied to clipboard

Challenge: Humans produce and interpret complex utterances even in simple scenarios.
Approach: They present a large-scale English language corpus with 34,268 (polar question, indirect answer) pairs to enable progress on this task.
Outcome: The proposed corpus contains 34,268 (polar question, indirect answer) pairs, and reaches 82-88% accuracy for a 4-class distinction, and 64-85% for 6 classes.
StyleKQC: A Style-Variant Paraphrase Corpus for Korean Questions and Commands (2022.lrec-1)

Copied to clipboard

Challenge: Especially for questions and commands, style-variant paraphrasing can be crucial in tone and manner.
Approach: They propose a corpus construction scheme that considers intent and formality of directives in Korean language.
Outcome: The proposed method is validated by a corpus construction scheme on Korean topics.
KC-GenRe: A Knowledge-constrained Generative Re-ranking Method Based on Large Language Models for Knowledge Graph Completion (2024.lrec-main)

Copied to clipboard

Challenge: Knowledge graph completion (KGC) is a critical task to predict missing facts among entities.
Approach: They propose a knowledge-constrained generative re-ranking method based on generative large language models for KGC that can predict missing facts among entities.
Outcome: The proposed method achieves state-of-the-art performance on four datasets and 9.0% and 11.1% compared to the previous methods.
Towards a Zero-Data, Controllable, Adaptive Dialog System (2024.lrec-main)

Copied to clipboard

Challenge: Recent approaches to controllable dialog systems require additional training data to be deployed in new domains.
Approach: They propose to generate dialog tree data directly from dialog trees by using a commercial Large Language Model or a single GPU.
Outcome: The proposed approach can achieve comparable dialog success to models trained on human data.
UniPCM: Universal Pre-trained Conversation Model with Task-aware Automatic Prompt (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies have shown that multi-task instruction tuning after pre-training greatly improves the model’s robustness and transfer ability, which is crucial for building a high-quality dialog system.
Approach: They propose to use Task-aware Automatic Prompt generation (TAP) to automatically generate high-quality prompts from 15 dialog-related tasks.
Outcome: The proposed model is robust to input prompts and capable of various dialog-related tasks.
Frame of Reference: Addressing the Challenges of Common Ground Representation in Situational Dialogs (2026.findings-acl)

Copied to clipboard

Challenge: Prior studies have demonstrated that Large Language Models (LLMs) are capable of performing grounding acts such as requesting clarification or producing acknowledgments, yet relatively little work has investigated how common ground can be explicitly represented and stored for later use.
Approach: They propose to use relational references to represent common ground in situated dialogues and propose to improve both the establishment of common ground and its subsequent use in the conversation.
Outcome: The proposed models can establish and exploit common ground in situated dialogues and improve its subsequent use.
DarwinTOD: LLM-Driven Lifelong Self-evolution for Task-oriented Dialog Systems (2026.acl-long)

Copied to clipboard

Challenge: Continual learning approaches fail to achieve autonomy lifelong improvement in dynamic environments . current task-oriented dialog systems are static, unable to learn from ongoing interactions .
Approach: They propose a lifelong self-evolving dialog framework that integrates evolutionary computation and LLM driven self-improvement into a single framework.
Outcome: The proposed framework surpasses state-of-the-art methods and exhibits continuous performance gains throughout evolution.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations